8 Parallel File Systems

نویسندگان

  • Robert Ross
  • Philip Carns
  • David Metheny
چکیده

The success of a CDI Grid is dependent upon the design of its storage infrastructure. As seen in Chapter 7, processing in this environment revolves around the simultaneous movement and transformation of data on many compute elements. Effective storage solutions combine hardware and software to meet these needs. The storage hardware selected must provide enough raw throughput for the expected workloads. Typical storage hardware architectures also often provide some redundancy to help in creating a fault tolerant system. Storage software, specifically file systems, must organize this storage hardware into a single logical space, provide efficient mechanisms for accessing that space, and hide common hardware failures from compute elements. Parallel file systems (PFSes) are a particular class of file systems that are well suited to this role. This chapter will describe a variety of PFS architectures, but the key feature that classifies all of them as parallel file systems is their ability to support true parallel I/O. Parallel I/O in this context means that many compute elements can read from or write to the same files concurrently without significant performance degradation and without data corruption. This is the critical element that differentiates PFSes from more traditional network file systems, and it is this characteristic that makes PFSes an integral part of an effective CDI Grid. While databases play a significant role in data engineering for the grid, they are not appropriate for all types of data. PFSes are particularly well-suited to storing large amounts of structured data that can reside within flat files. This type of data can be dynamically distributed and processed across almost any number of compute elements by simply assigning each compute element a unique portion of the file. Compute elements can then be added incrementally to increase the processing rate. This corresponds well with CDI processes that will be discussed in greater

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Project Report for Project Shared Parallel File System

In recent years, shared parallel file system has been a new hot area for the high throughput computing. Many systems have been developed for this purpose, which include PVFS2 and GFS. These systems were being studied in my project. In this project, the concept of such a shared parallel file system has been built up by installing it on four nodes of the LISA cluster at SARA (Stichting Academisch...

متن کامل

Evaluating Algorithms for Shared File Pointer Operations in MPI I/O

MPI-I/O is a part of the MPI-2 specification defining file I/O operations for parallel MPI applications. Compared to regular POSIX style I/O functions, MPI I/O offers features like the distinction between individual file pointers on a per-process basis and a shared file pointer across a group of processes. The objective of this study is the evaluation of various algorithms of shared file pointe...

متن کامل

vPFS: Bandwidth Virtualization of Parallel Storage Systems

This paper presents vPFS, a new parallel file system performance management approach to support the allocation of shared storage bandwidth on a per-application basis. Existing parallel file systems are unable to differentiate I/O requests from different applications and meet per-application bandwidth requirements. This limitation presents an increasing hurdle for applications to achieve their d...

متن کامل

U.S. Department of Energy Best Practices Workshop on File Systems & Archives: Usability at Los Alamos National Lab

There yet exist no truly parallel file systems. Those that make the claim fall short when it comes to providing adequate concurrent write performance at large scale. This limitation causes large usability headaches in HPC computing. Users need two major capabilities missing from current parallel file systems. One, they need low latency interactivity. Two, they need high bandwidth for large para...

متن کامل

Data - intensive file systems for Internet services : A rose by any other

Data-intensive distributed file systems are emerging as a key component of large scale Internet services and cloud computing platforms. They are designed from the ground up and are tuned for specific application workloads. Leading examples, such as the Google File System, Hadoop distributed file system (HDFS) and Amazon S3, are defining this new purpose-built paradigm. It is tempting to classif...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2009